Papers with sequence-to-sequence model

68 papers
DART: A Lightweight Quality-Suggestive Data-to-Text Annotation Tool (2020.coling-demos)

Copied to clipboard

Challenge: Neural data-to-text generation systems require large-scale labeled data to generate sentences.
Approach: They propose to create an interactive annotation tool that iteratively analyzes annotated structured data to better sample unlabeled data.
Outcome: The proposed tool reduces the number of annotations needed with active learning and automatically suggests relevant labels.
BME-UW at SRST-2019: Surface realization with Interpreted Regular Tree Grammars (D19-63)

Copied to clipboard

Challenge: adaamko's system restores word order and inflection from a graph of typed, directed dependencies between lemmas.
Approach: They propose a method that restores word order and inflection from a graph of typed, directed dependencies between lemmas.
Outcome: The proposed system restores word order and inflection from a graph of typed, directed dependencies between lemmas.
Deep Bayesian Natural Language Processing (P19-4)

Copied to clipboard

Challenge: Introduction to deep Bayesian learning for natural language addresses the fundamentals of statistical models and neural networks.
Approach: This tutorial addresses the advances in deep Bayesian learning for natural language . it focuses on advanced Bayessian models and deep models . authors present case studies and domain applications to tackle different issues .
Outcome: This tutorial focuses on advanced Bayesian models and deep models for natural language . case studies and domain applications are presented to tackle different issues in deep Bayessian processing, learning and understanding.
Interactive Plot Manipulation using Natural Language (2021.naacl-demos)

Copied to clipboard

Challenge: a new interactive plotting agent is available for programming with natural language . the interactive aspect allows users to manipulate plots using natural language instructions.
Approach: They propose an interactive natural language interface for plotting that maps language to plot updates.
Outcome: The proposed system maps language to plot updates within an interactive programming environment.
Break It Down: A Question Understanding Benchmark (2020.tacl-1)

Copied to clipboard

Challenge: Understanding natural language questions entails the ability to break down a question into the requisite steps for computing its answer.
Approach: They introduce a Question Decomposition Meaning Representation (QDMR) for questions . they demonstrate that QDMRs can be annotated at scale using a hotpotQA dataset .
Outcome: The proposed model outperforms several natural baselines in the open-domain question answering hotpotQA dataset and can be deterministically converted to a pseudo-SQL formal language.
Answer-based Adversarial Training for Generating Clarification Questions (N19-1)

Copied to clipboard

Challenge: a goal of natural language processing is to develop techniques that enable machines to process naturally occurring language.
Approach: They propose a model where hypothetical answers are latent variables that can guide the model into generating more useful clarification questions.
Outcome: The proposed model outperforms retrieval-based models and ablations that exclude utility model and adversarial training on two datasets.
Extract, Transform and Filling: A Pipeline Model for Question Paraphrasing based on Template (D19-55)

Copied to clipboard

Challenge: Recent approaches for paraphrasing generate unpredictable results .
Approach: They propose a question paraphrasing pipeline model based on templates that identifies template and retrieves candidate templates and fills them with original topic words.
Outcome: The proposed model outperforms the seq2seq model on two datasets and is more promising when the training sample is small.
Cool English: a Grammatical Error Correction System Based on Large Learner Corpora (C18-2)

Copied to clipboard

Challenge: Existing systems that correct grammatical errors are lacking in second language learning due to limited vocabulary and inadequate command of grammar.
Approach: They propose a grammatical error correction system that provides corrective feedback for essays using a sequence-to-sequence model.
Outcome: The proposed system achieves competitive performance on a number of publicly available testsets.
Neural Math Word Problem Solver with Reinforcement Learning (C18-1)

Copied to clipboard

Challenge: Existing models for solving math word problems rely on predefined rules or feature engineering.
Approach: They propose to incorporate copy and alignment mechanism into the sequence-to-sequence model to address two shortcomings . they use model output as a feature and incorporate it into the feature-based model to explore the effectiveness .
Outcome: The proposed model outperforms the state-of-the-art models on the problem solving task.
Query and Output: Generating Words by Querying Distributed Word Representations for Paraphrase Generation (N18-1)

Copied to clipboard

Challenge: Existing models tend to memorize words instead of learning meaning of words . existing models tend not to model semantic information, resulting in incorrect sentences .
Approach: They propose a novel model that generates words by querying distributed word representations . they evaluate model on two paraphrase-oriented tasks, namely text simplification and short abstractive summarization .
Outcome: The proposed model outperforms the baseline model on two paraphrase-oriented tasks . it achieves state-of-the-art performance on these benchmark datasets .
Automatically Summarizing Evidence from Clinical Trials: A Prototype Highlighting Current Challenges (2023.eacl-demo)

Copied to clipboard

Challenge: Existing systems that retrieve trial publications matching a query are inefficient and introduce unsupported statements.
Approach: They propose a system that aims to automatically summarize evidence presented in the set of randomized controlled trials most relevant to a given query.
Outcome: The proposed system retrieves trial publications matching a query specifying a combination of condition, intervention(s), and outcome(s) and ranks them according to sample size and estimated study quality.
Empirical Error Modeling Improves Robustness of Noisy Neural Sequence Labeling (2021.findings-acl)

Copied to clipboard

Challenge: Standard sequence labeling systems fail when processing noisy user-generated text or consuming the output of an OCR process.
Approach: They propose an empirical error generation approach that employs a sequence-to-sequence model trained to perform translation from error-free to erroneous text.
Outcome: The proposed method outperforms baseline noise generation and error correction techniques on the erroneous sequence labeling data sets.
Learning the Extraction Order of Multiple Relational Facts in a Sentence with Reinforcement Learning (D19-1)

Copied to clipboard

Challenge: Existing works didn’t consider the extraction order of relational facts in a sentence.
Approach: They propose to take the extraction order into consideration by applying reinforcement learning into a sequence-to-sequence model.
Outcome: The proposed model could generate relational facts freely.
Task-Oriented Dialogue as Dataflow Synthesis (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches to task-oriented dialogue represent dialogue state as a dataflow graph . microsoft's SMCalFlow dataset features complex dialogues about events, weather, places, and people .
Approach: They propose a dataflow graph-based dialogue agent that maps each user utterance to a program that extends this graph.
Outcome: The proposed framework improves representability and predictability in natural dialogues . it uses dataflow graphs and metacomputation to map user intents to a program .
Using Intermediate Representations to Solve Math Word Problems (P18-1)

Copied to clipboard

Challenge: Existing approaches to solving math word problems do not include higher-order operations that cannot be explicitly represented in equations.
Approach: They propose an iterative labeling framework that generates intermediate forms and executes them to obtain the final answers.
Outcome: The proposed model outperforms existing models in solving math word problems.
Multilingual Denoising Pre-training for Neural Machine Translation (2020.tacl-1)

Copied to clipboard

Challenge: Existing approaches to pre-train models focus on only English corpora, but this is not common in machine translation.
Approach: They propose a sequence-to-sequence denoising auto-encoder pre-trained on monolingual corpora . they show that it produces significant performance gains across MT tasks .
Outcome: The proposed model can achieve significant performance gains across a wide variety of MT tasks.
Using Question Answering Rewards to Improve Abstractive Summarization (2021.findings-emnlp)

Copied to clipboard

Challenge: Neural abstractive summarization models have seen improvements in recent years, but they still suffer from multiple drawbacks.
Approach: They propose a general framework to train abstractive summarization models to alleviate these issues by question-answering based rewards.
Outcome: The proposed framework is preferred over general abstractive summarization models.
Document Ranking with a Pretrained Sequence-to-Sequence Model (2020.findings-emnlp)

Copied to clipboard

Challenge: Experimental results on the MS MARCO passage ranking task show that our ranking approach is superior to strong encoder-only models.
Approach: They propose to use a pretrained sequence-to-sequence model to generate relevance labels as "target tokens" they also show how the underlying logits of these target tokens can be interpreted as relevance probabilities for ranking.
Outcome: The proposed model outperforms existing models in a data-poor setting and significantly outperformed an encoder-only model on the MS MARCO passage ranking task.
Neural Data-to-Text Generation with LM-based Text Augmentation (2021.eacl-main)

Copied to clipboard

Challenge: Neural data-to-text generation is a difficult task for many new applications because of a lack of training data.
Approach: They propose a few-shot approach that augments the data available for training by generating new text samples based on replacing specific values by alternative ones from the same category and pairing the new text with data samples.
Outcome: The proposed approach outperforms fully supervised sequence-to-sequence models with less than 10% of the training set on both datasets.
Entity Commonsense Representation for Neural Abstractive Summarization (N18-1)

Copied to clipboard

Challenge: Current ELS’s are not sufficiently effective, possibly introducing unresolved ambiguities and irrelevant entities.
Approach: They propose an off-the-shelf entity linking system to extract linked entities and propose Entity2Topic (E2T) module attachable to a sequence-to-sequence model that transforms a list of entities into a vector representation of the topic of the summary.
Outcome: The proposed model improves the performance of the Gigaword and CNN summarization datasets by at least 2 ROUGE points.
ParsTranslit: Truly Versatile Tajik-Farsi Transliteration (2026.findings-eacl)

Copied to clipboard

Challenge: Despite significant similarities between the two written standards, script differences hinder simple one-to-one mapping, hindering written communication and interaction between Tajikistan and its Persian-speaking “siblings”.
Approach: They propose to use a sequence-to-sequence model to convert between two scripts in a Persian-speaking country using two datasets.
Outcome: The proposed model achieves chrF++ and Normalized CER scores of 87.91 and 0.05 from Farsi to Tajik and 92.28 and 0.04 from Tajikistan to Farsis.
Compositional Generalization and Natural Language Variation: Can a Semantic Parsing Approach Handle Both? (2021.acl-long)

Copied to clipboard

Challenge: Existing approaches to semantic parsing only evaluated on synthetic datasets that are not representative of natural language variation.
Approach: They propose a semantic parsing approach that handles both natural language variation and compositional generalization.
Outcome: The proposed model outperforms existing models across compositional generalization challenges on non-synthetic datasets while being competitive with the state-of-the-art on standard evaluations.
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation (2025.coling-main)

Copied to clipboard

Challenge: Pre-trained sequence-to-sequence models are typically pretrained on extensive raw text corpora and fine-tuned on task-specific data.
Approach: They introduce a pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora.
Outcome: The proposed model outperforms existing models on three generative tasks and is data-efficient and effective in enhancing performance across various natural language generation tasks.
Extracting Symptoms and their Status from Clinical Conversations (P19-1)

Copied to clipboard

Challenge: Existing models for extracting symptoms from clinical conversations are inherently difficult.
Approach: They propose two new deep learning models tailored for a new application . they propose a hierarchical span-attribute tagging model and a sequence-to-sequence model .
Outcome: The proposed models perform well under different conditions and are compared to existing models.
Compositional Generalization via Semantic Tagging (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing neural sequence-to-sequence models fail at compositional generalization, i.e., they cannot generalize to unseen compositions of seen components.
Approach: They propose a decoding framework that preserves expressivity and generality of sequence-to-sequence models while featuring lexicon-style alignments and disentangled information processing.
Outcome: The proposed framework improves compositional generalization across model architectures, domains, and semantic formalisms on three semantic parsing datasets.
An Empirical Study of Building a Strong Baseline for Constituency Parsing (P18-2)

Copied to clipboard

Challenge: Sequence-to-sequence models have been used for natural language generation tasks such as machine translation and summarization.
Approach: They propose to build a strong baseline based on general purpose sequence-to-sequence models for constituency parsing.
Outcome: The proposed model outperforms existing models in natural language generation tasks without any explicit task-specific knowledge or architecture of constituent parsing.
Utilizing Character and Word Embeddings for Text Normalization with Sequence-to-Sequence Models (D18-1)

Copied to clipboard

Challenge: Recent advances in text normalization have limited applications in other languages . a novel approach to text normalizing uses character embeddings and word embedds .
Approach: They propose a sequence-to-sequence model with character-based attention that uses pre-trained word embeddings to model subword information.
Outcome: The proposed model achieves state-of-the-art F1 score on Arabic spelling correction task despite being small and unsuited for the task.
Learning to Control the Specificity in Neural Response Generation (P18-1)

Copied to clipboard

Challenge: Existing generative conversational models tend to favor general and trivial responses which appear frequently.
Approach: They propose a controlled response generation mechanism to handle different utterance-response relationships in terms of specificity.
Outcome: The proposed model outperforms state-of-the-art models under automatic and human evaluations.
Autoencoder as Assistant Supervisor: Improving Text Representation for Chinese Social Media Text Summarization (P18-2)

Copied to clipboard

Challenge: Existing abstractive text summarization models learn a semantic representation of the source text and the summaries from it.
Approach: They evaluate the model on a popular Chinese social media dataset and compare it to other models.
Outcome: The proposed model achieves state-of-the-art performance on a popular Chinese social media dataset.
On the Helpfulness of Document Context to Sentence Simplification (2020.coling-main)

Copied to clipboard

Challenge: Text simplification is a hot issue in the field of natural language generation (NLG).
Approach: They propose to use Wikipedia context to improve sentence simplification by using neural networks to learn the effects of preceding and following sentences on current sentences.
Outcome: The proposed model outperforms the best performing model on the baseline dataset by 2.46 (7.22%).
mT6: Multilingual Pretrained Text-to-Text Transformer with Translation Pairs (2021.emnlp-main)

Copied to clipboard

Challenge: Multilingual T5 pretrains a sequence-to-sequence model on monolingual texts, but it has shown promising results on many cross-lingual tasks.
Approach: They propose a partially non-autoregressive objective for text-to-text pre-training and propose mT6 to improve cross-lingual transferability over multilingual T5.
Outcome: The proposed model improves cross-lingual transferability over existing models.
Jointly Optimizing Diversity and Relevance in Neural Response Generation (N19-1)

Copied to clipboard

Challenge: Recent neural conversation models often generate bland and generic responses . however, the improvement often comes at the cost of decreased relevance .
Approach: They propose a spacefusion model to jointly optimize diversity and relevance that fuses the latent space of a sequence-to-sequence model and that of an autoencoder model by leveraging novel regularization terms.
Outcome: The proposed model improves diversity and relevance compared to baselines in both diversity and diversity.
SeaD: End-to-end Text-to-SQL Generation with Schema-aware Denoising (2022.findings-naacl)

Copied to clipboard

Challenge: Using sketch-based slot filling, text-to-SQL models suffer from over-complexity . et al., e.al., and d.albert, dr., propose a novel method for text- to-Sql generation .
Approach: They propose to train sequence-to-sequence model with Schema-aware Denoising . they propose a clause-sensitive execution guided (EG) decoding strategy .
Outcome: The proposed method improves performance in schema linking and grammar correctness . it also establishes new state-of-the-art on the WikiSQL benchmark .
A Graph-to-Sequence Model for AMR-to-Text Generation (P18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) is a semantic formalism that encodes the meaning of a sentence as a rooted, directed graph.
Approach: They propose a neural graph-to-sequence model that leverages LSTM to encode a linearized AMR structure.
Outcome: The proposed model outperforms existing methods on a benchmark.
Generative Knowledge Selection for Knowledge-Grounded Dialogues (2023.findings-eacl)

Copied to clipboard

Challenge: Knowledge selection is the key in knowledge-grounded dialogues (KGD), which aims to select an appropriate knowledge snippet to be used in the utterance based on dialogue history.
Approach: They propose a generative approach for knowledge selection called GenKS that learns to select snippets by generating their identifiers with a sequence-to-sequence model.
Outcome: The proposed approach captures intra-knowledge interaction inherently through attention mechanisms while generating their identifiers with a sequence-to-sequence model.
A Weak Supervision Approach for Few-Shot Aspect Based Sentiment Analysis (2024.eacl-long)

Copied to clipboard

Challenge: Existing methods to improve few-shot performance in aspect-based sentiment analysis (ABSA) require complex interactions between the target and the polarity of the sentiment.
Approach: They propose a pipeline approach to construct a noisy ABSA dataset and adapt it to the ABSA tasks.
Outcome: The proposed model outperforms the state-of-the-art on the aspect extraction sentiment classification task and is capable of performing the harder aspect sentiment triplet extraction task.
Adversarial Attack and Defense of Structured Prediction Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to building effective adversarial attackers focus on classification problems.
Approach: They propose a framework that learns to attack a structured prediction model with feedbacks from multiple reference models.
Outcome: The proposed framework is able to attack state-of-the-art models and boost them with training . it is based on a sequence-to-sequence model with feedbacks from multiple reference models .
DyKgChat: Benchmarking Dialogue Generation Grounding on Dynamic Knowledge Graphs (D19-1)

Copied to clipboard

Challenge: Existing work has not shown that knowledge-grounded models can zero-shot adapt to updated, unseen knowledge graphs.
Approach: They propose a task to apply dynamic knowledge graphs to neural conversation models . they propose 'dyKgChat' that selects an output from two networks at each time step .
Outcome: The proposed model outperforms existing knowledge-grounded conversation models in evaluation metrics.
Automatic Keyphrase Generation by Incorporating Dual Copy Mechanisms in Sequence-to-Sequence Learning (2022.coling-1)

Copied to clipboard

Challenge: Existing models for keyphrase generation use a copy mechanism to generate keyphrases, but they do not identify key words in the source text and copy them to create more keyphrase.
Approach: They propose a dual-copier keyphrase generation model that uses a sequence-to-sequence model to generate keyphrases for a piece of text.
Outcome: The proposed model outperforms baseline models and achieves an obvious performance improvement.
Unified Pre-training for Program Understanding and Generation (2021.naacl-main)

Copied to clipboard

Challenge: PLUG is a programming language that is used for programming and language understanding and generation tasks.
Approach: They propose a sequence-to-sequence model that performs a broad spectrum of program and language understanding and generation tasks.
Outcome: The proposed model outperforms or rivals state-of-the-art models on code summarization, code generation, and code translation tasks in seven programming languages.
ProphetNet: Predicting Future N-gram for Sequence-to-SequencePre-training (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing sequence-to-sequence models are optimized for future n-gram prediction and n stream self-attention mechanism.
Approach: They propose a self-supervised objective called future n-gram prediction and the proposed n stream self-attention mechanism to optimize the model for sequence-to-sequence learning.
Outcome: The proposed model achieves state-of-the-art on CNN/DailyMail, Gigaword, and SQuAD 1.1 benchmarks compared to the models using the same scale pre-training corpus.
Multi-Input Attention for Unsupervised OCR Correction (P18-1)

Copied to clipboard

Challenge: Existing methods for OCR correction are mostly supervised methods that correct recognition errors in a single output.
Approach: They propose a sequence-to-sequence model with attention and a decoder with attention averaging to search for consensus among multiple sequences.
Outcome: The proposed methods cut the character and word error rates nearly in half on single inputs and can rival supervised methods.
MuCGEC: a Multi-Reference Multi-Source Evaluation Dataset for Chinese Grammatical Error Correction (2022.naacl-main)

Copied to clipboard

Challenge: Using a multi-reference multi-source evaluation dataset, Chinese grammatical error correction (CGEC) is relatively scarce.
Approach: They propose a multi-reference multi-source evaluation dataset for Chinese grammar error correction . the dataset contains 7,063 sentences written by Chinese-as-a-Second-Language learners .
Outcome: The proposed dataset can be used to evaluate Chinese grammar errors in Chinese.
Neural Text Generation from Rich Semantic Representations (N19-1)

Copied to clipboard

Challenge: 2 is a neural model that maps a linearization of Dependency MRS to text . 1 is based on a BLEU score of 66.11 when trained on gold data .
Approach: They propose to use Minimal Recursion Semantics to generate high-quality text from structured representations.
Outcome: The proposed model achieves a BLEU score of 77.17 on the full test set and 83.37 on the subset of test data most closely matching the silver data domain.
Generative Emotion Cause Triplet Extraction in Conversations with Commonsense Knowledge (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on ECTEC focus on Causal Emotion Entailment and Emotion-Cause Pair Extraction in Conversations.
Approach: They propose to decompose the ECTEC task into multiple subtasks and solve them in a pipeline manner.
Outcome: The proposed model outperforms competing systems on two benchmark datasets.
Fluent Translations from Disfluent Speech in End-to-End Speech Translation (N19-1)

Copied to clipboard

Challenge: Disfluency removal is an intermediate step between speech recognition and machine translation (MT) with the rise of end-to-end speech translation systems, disfluency recognition and removal needs to be incorporated into the model architectures or handled as a post-processing step.
Approach: They propose to use a sequence-to-sequence model to translate from noisy, disfluent speech to fluent text with disfluencies removed using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset.
Outcome: The proposed model generates fluent translations from disfluent speech using the recently collected ‘copy-edited’ references for the Fisher Spanish-English dataset.
Mapping probability word problems to executable representations (2021.emnlp-main)

Copied to clipboard

Challenge: a recent paper addresses the problem of solving math word problems automatically . a number of approaches have been proposed for solving word problems .
Approach: They employ a sequence-to-sequence model to generate intermediate representations for word problems . they then use a probabilistic programming system to provide the answer . their best performing model incorporates general-domain contextualised word representations .
Outcome: The proposed model is the best performing on a declarative language and a probabilistic programming system.
IMaT: Unsupervised Text Attribute Transfer via Iterative Matching and Translation (D19-1)

Copied to clipboard

Challenge: Existing approaches to rewrite sentences with certain attributes are difficult and often result in poor content-preservation and ungrammaticality.
Approach: They propose a method that uses a sequence-to-sequence model to learn attribute transfer . existing approaches try to explicitly disentangle content and attribute information .
Outcome: The proposed method outperforms complex state-of-the-art systems by a large margin in sentiment modification and formality transfer tasks.
Chinese Pinyin Aided IME, Input What You Have Not Keystroked Yet (D18-1)

Copied to clipboard

Challenge: Chinese pinyin input method engine (IME) converts pinyine into character based on its core component, pinyan-to-character conversion (P2C).
Approach: They propose a sequence-to-sequence model with gated-attention mechanism for Chinese IMEs.
Outcome: The proposed model improves on existing models in benchmark datasets showing great user experience improvement compared to traditional models.
Improving Compositional Generalization with Latent Structure and Data Augmentation (2022.naacl-main)

Copied to clipboard

Challenge: Generic unstructured neural networks struggle on out-of-distribution compositional generalization.
Approach: They propose a method to recombinate examples from a model called Compositional Structure Learner and add them to a pre-trained sequence-to-sequence model.
Outcome: The proposed model is even stronger than a T5-CSL ensemble on two real world compositional generalization tasks.
Contextualizing Generated Citation Texts (2024.lrec-main)

Copied to clipboard

Challenge: Abstractive citation text generation is usually framed as an infilling task . however, examining a recent LED-based citation generation system, we find that many of the generated citations are generic summaries of the reference paper’s main contribution, ignoring the citation context’s focus on a different topic.
Approach: They propose a modification to the citation text generation task by training the generation model to generate a citation given a reference paper and the context window around the target.
Outcome: The proposed model can generate citations based on the entire context window, including the target citation.
Controllable Paraphrase Generation for Semantic and Lexical Similarities (2024.lrec-main)

Copied to clipboard

Challenge: Lexically diverse paraphrases are crucial in data augmentation because they enhance the linguistic diversity of the corpus.
Approach: They propose a controllable model for semantic and lexical similarities by attaching tags to the head of the input sentence.
Outcome: The proposed model can paraphrase an input sentence according to the tags specified.
DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion (N19-1)

Copied to clipboard

Challenge: Existing datasets for sentence fusion are small and insufficient for training modern neural models.
Approach: They propose a method for automatically-generating fusion examples from raw text . they apply their method to Wikipedia and Sports articles to generate fusion models .
Outcome: The proposed method improves performance on WebSplit when viewed as a sentence fusion task.
Linguistically-Informed Specificity and Semantic Plausibility for Dialogue Generation (N19-1)

Copied to clipboard

Challenge: Past work has focused on word frequency-based approaches to improving specificity, such as penalizing responses with only common words.
Approach: They propose to rerank a sequence-to-sequence model to improve the informativeness, reasonableness, and grammatically of responses by using externally-trained classifiers targeting each of these factors.
Outcome: The proposed model improves the informativeness, reasonableness, and grammatically of responses.
Generic resources are what you need: Style transfer tasks without task-specific parallel training data (2021.emnlp-main)

Copied to clipboard

Challenge: Text style transfer is a task aimed at converting a text of one style into another while preserving its content.
Approach: They propose a multi-step procedure which builds on a generic pre-trained sequence-to-sequence model and an iterative back-translation approach to train two models in a transfer direction.
Outcome: The proposed method outperforms existing unsupervised approaches on the two most popular style transfer tasks: formality transfer and polarity swap.
Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT (2020.emnlp-main)

Copied to clipboard

Challenge: a new method of analysis based on semantic tags demonstrates that character-level representations improve performance across a subset of selected semantic phenomena.
Approach: They combine character-level and contextual language model representations to improve performance on Discourse Representation Structure parsing.
Outcome: The proposed model improves performance on a subset of selected semantic phenomena.
Consistent Response Generation with Controlled Specificity (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to generate fluent responses generate inconsistent responses . we use a sequence-to-sequence model to generate specific responses based on a co-occurrence degree .
Approach: They propose a method to control the specificity of responses while maintaining the consistency with the utterances.
Outcome: The proposed method produces highly consistent responses in open-domain dialogues . it can generate fluent responses while maintaining the consistency with the utterances compared to the conventional model .
Tackling the Low-resource Challenge for Canonical Segmentation (2020.emnlp-main)

Copied to clipboard

Challenge: morphological segmentation is a task of dividing words into their constituting morphemes . we compare two new approaches for the task when training data is limited .
Approach: They propose to use an LSTM pointer-generator and a sequence-to-sequence model to perform canonical segmentation when training data is limited.
Outcome: The proposed models outperform existing models on German, English, and Indonesian in low-resource scenarios by 11.4% accuracy.
Document-level Entity-based Extraction as Template Generation (2021.emnlp-main)

Copied to clipboard

Challenge: Document-level entity-based extraction (EE) tasks extract entity-centric information from unstructured text across multiple sentences.
Approach: They propose a generative framework for two document-level EE tasks: role-filler entity extraction (RE) and relation extraction ( RE).
Outcome: The proposed framework captures cross-entity dependencies and avoids exponential computation complexity of identifying N-ary relations.
Answer-focused and Position-aware Neural Question Generation (D18-1)

Copied to clipboard

Challenge: Recent neural network-based approaches generate interrogative words that do not match the answer type.
Approach: They propose an answer-focused and position-aware neural question generation model to address these issues.
Outcome: The proposed model outperforms the baseline and outperformed the state-of-the-art system.
Do RNN States Encode Abstract Phonological Alternations? (2021.naacl-main)

Copied to clipboard

Challenge: Sequence-to-sequence models have been successful in word formation tasks, but the opacity of the models makes it difficult to determine whether complex generalizations are learned or whether there is some level of generalization across related sound changes.
Approach: They propose to train character-based sequence-to-sequence models for inflection of Finnish nouns into the genitive case, an inflation type which is encoded in the hidden states of an LSTM encoderdecoder trained to perform word infference.
Outcome: The proposed models encode 17 different consonant gradation processes in a handful of dimensions in the RNN.
Universal Conditional Masked Language Pre-training for Neural Machine Translation (2022.acl-long)

Copied to clipboard

Challenge: Pre-trained sequence-to-sequence models have significantly improved Neural Machine Translation (NMT) this paper demonstrates that pre-training a sequence- to-squence model with a bidirectional decoder can produce notable performance gains for both Autoregressive and Non-autoregressive NMT tasks.
Approach: They propose a conditional masked language model pre-trained on bilingual and monolingual corpora in many languages.
Outcome: The proposed model can achieve significant performance improvements on all scenarios from low- to extremely high-resource languages.
Grounded Multimodal Named Entity Recognition on Social Media (2023.acl-long)

Copied to clipboard

Challenge: Existing studies on Multimodal Named Entity Recognition only extract entity-type pairs in text, which is useless for multimodal knowledge graph construction.
Approach: They propose a task to identify named entities in text and their bounding box groundings in image . they extend four well-known MNER methods to establish a number of baseline systems .
Outcome: The proposed framework outperforms baseline systems on the GMNER task.
DiffusionRet: Diffusion-Enhanced Generative Retriever using Constrained Decoding (2023.findings-emnlp)

Copied to clipboard

Challenge: Generative retrieval methods have suffered from the lack of the intermediate reasoning step . generative retrieval uses sequence-to-sequence diffusion models to map a query to relevant docids .
Approach: They propose a novel method that uses query as an intermediate step before retrieval . they propose to use sequence-to-sequence diffusion models to map a query to relevant docids .
Outcome: Experiments show that proposed method outperforms existing methods on MARCO and Natural Questions datasets.
Predicting generalization performance with correctness discriminators (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models estimate accuracy of models on unlabeled test data, but they hide their own uncertainty.
Approach: They propose a model that establishes upper and lower bounds on the accuracy without requiring gold labels for the unseen data.
Outcome: The proposed model establishes upper and lower bounds on accuracy without requiring gold labels for the unseen data.
Low-resource Neural Machine Translation: Benchmarking State-of-the-art Transformer for Wolof<->French (2022.lrec-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) systems can translate between French (FR) 1 and Wolof (WO, ISO 639-3), a lowresource Niger-Congo language mainly spoken in Senegal (Gamble, 1950).
Approach: They propose two neural machine translation systems based on sequence-to-sequence with attention and Transformer architectures to translate between French (FR) 1 and Wolof (WO, ISO 639-3).
Outcome: The proposed models outperform the classic sequence-to-sequence model in all settings and are less sensitive to noise.
JADE: Corpus for Japanese Definition Modelling (2022.lrec-1)

Copied to clipboard

Challenge: Existing corpus for definition modelling techniques is limited to English . this study aimed to develop a corpus that provides definitions of words and phrases .
Approach: They investigated and released a corpus for Japanese definition modelling . the JADE provides 630k sets of targets, their definitions, and usage examples as contexts .
Outcome: The JADE corpus provides 630k sets of targets, their definitions, and usage examples as contexts for 41k unique targets.
Generate then Refine: Data Augmentation for Zero-shot Intent Detection (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing data augmentation methods rely on few labelled examples for each intent category, which can be expensive in settings with many possible intents.
Approach: They propose a data augmentation method for intent detection in zero-resource domains by using an open-source large language model and a smaller sequence-to-sequence model.
Outcome: The proposed method significantly improves the data utility and diversity over the zero-shot LLM baseline for unseen domains and over common baseline approaches.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations